Papers with multilingual baselines

3 papers
Paraphrases as Foreign Languages in Multilingual Neural Machine Translation (P19-2)

Copied to clipboard

Challenge: Unlike previous studies that use paraphrases at the word/phrase level, we train on parallel paraphrase training on closely related languages.
Approach: They train on parallel paraphrases in the style of multilingual Neural Machine Translation (NMT) they train on translations of the whole corpus that are consistent in structure as paraphrase versions at the corpus level.
Outcome: The proposed training on paraphrases outperforms the baselines on two languages and improves lexical choice and entropy.
IT5: Text-to-text Pretraining for Italian Language Understanding and Generation (2024.lrec-main)

Copied to clipboard

Challenge: Xue et al., 2022) use the text-to-text paradigm to train multilingual models.
Approach: They introduce the first family of encoder-decoder transformer models pretrain specifically on Italian and introduce the ItaGen benchmark to evaluate the models' performance.
Outcome: The proposed model outperforms models with multilingual baselines and the original model on English data.
Targeted Multilingual Adaptation for Low-resource Language Families (2024.findings-emnlp)

Copied to clipboard

Challenge: Massively multilingual models are known to have limited utility in any one language, and to perform poorly on low-resource languages.
Approach: They propose to adapt a pre-trained multilingual model to a language family and evaluate its performance on two downstream tasks and 11 evaluation languages.
Outcome: The proposed model outperforms mono- and multilingual models on two downstream tasks and 11 evaluation languages.

What is GenGO?

GenGO is an NLP powered publication search system. It currenctly indexes 30k+ papers from ACL Anthology, and implements multi-aspect summarization, semantic search, and more!

Information

About
Limitations